Papers with lightweight VLMs
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks (2026.findings-acl)
Copied to clipboard
Yuanze Hu, Xinyu Wang, Zhichao Yang, Gen Li, Ye Qiu, Zhaoxin Fan, Yifan Sun, Wenjun wu, Jin Dong, Xiaotie Deng
| Challenge: | Lightweight Vision-Language Models (VLMs) are indispensable for resource-constrained applications. |
| Approach: | They propose a framework that retrieves context from a memory bank to enhance alignment . they propose EMI-based approach to align vision and language models . |
| Outcome: | The proposed framework reduces training loss, accelerates convergence, and enhances task performance with negligible computational overhead. |
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding (2026.findings-eacl)
Copied to clipboard
| Challenge: | figurative language is essential for expressing intent, emotion, and perspective . figural language is often dependent on Styles Reasoning, causing incongruities between expressions . |
| Approach: | They propose a framework that induces reasoning capabilities to compact vision–language models . figurative language is essential in expressing intent, emotion, and perspective . |
| Outcome: | The proposed framework can interpret multimodal figurative language, provide transparent reasoning traces, and generalize across multiple figurativ styles. |
EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning (2023.findings-acl)
Copied to clipboard
| Challenge: | Pre-trained vision-language models have achieved impressive results in a range of vision-linguistic tasks. |
| Approach: | They propose a distilling then pruning framework to compress large vision-language models into smaller, faster ones. |
| Outcome: | The proposed framework reduces the size of a pre-trained large vision-language model and improves its performance on vision-linguistic tasks. |